Papers with Web-scale training
ICC : Quantifying Image Caption Concreteness for Multimodal Dataset Curation (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to curation text-image data are noisy and lack the fine-grained ability to isolate the most concrete samples that provide the strongest signal for learning in a noisy dataset. |
| Approach: | They propose a metric that evaluates caption text without an image reference to measure its concreteness and relevancy. |
| Outcome: | The proposed method detects the concreteness of captions without an image reference and correlates with human evaluation of concreteness in both single-word and caption-level texts. |